Automatic Text Summarization using Document Clustering Named Entity Recognition

نویسندگان

چکیده

Due to the rapid development of internet technology, social media and popular research article databases have generated many open text information. This large amount textual information leads 'Big Data'. Textual can be recorded repeatedly about an event or topic on different websites. Text summarization (TS) is emerging field that helps produce summary from a single multiple documents. The redundant in documents difficult, hence part all sentences may omitted without changing gist document. TS organized as exposition collect accents its special position, rather than being semantic nature. Non-ASCII characters pronunciation, including tokenizing lemmatization are involved generating summary. work has proposed Entity Aware Summarization using Document Clustering (EASDC) technique extract multi-documents. Named Recognition (NER) vital work. topics key terms identified NER technique. Extracted entities ranked with Zipf’s law sentence clusters formed k-means clustering. Cosine similarity-based used eliminate similar multi-documents unique EASDC evaluated CNN dataset it shown improvement 1.6 percentage when compared baseline methods Textrank Lexrank.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Named Entity Recognition in Persian Text using Deep Learning

Named entities recognition is a fundamental task in the field of natural language processing. It is also known as a subset of information extraction. The process of recognizing named entities aims at finding proper nouns in the text and classifying them into predetermined classes such as names of people, organizations, and places. In this paper, we propose a named entity recognizer which benefi...

متن کامل

Named Entity Recognition Using Web Document Corpus

This paper introduces a named entity recognition approach in textual corpus. This Named Entity (NE) can be a named: location, person, organization, date, time, etc., characterized by instances. A NE is found in texts accompanied by contexts: words that are left or right of the NE. The work mainly aims at identifying contexts inducing the NE’s nature. As such, The occurrence of the word "Preside...

متن کامل

Document Clustering and Text Summarization

This paper describes a text mining tool that performs two tasks, namely document clustering and text summarization. These tasks have, of course, their corresponding counterpart in “conventional” data mining. However, the textual, unstructured nature of documents makes these two text mining tasks considerably more difficult than their data mining counterparts. In our system document clustering i...

متن کامل

Automatic Text Summarization Using Lexical Clustering

The goal of automatic text summarization is to reduce the size of a document while preserving its content. We investigate a summarization method which uses not only statistical features but also the contextual meaning of documents by using lexical clustering. We present a new method to compute lexical cluster in a text without high cost knowledge resources; the WordNet thesaurus. Summarization ...

متن کامل

A survey on Automatic Text Summarization

Text summarization endeavors to produce a summary version of a text, while maintaining the original ideas. The textual content on the web, in particular, is growing at an exponential rate. The ability to decipher through such massive amount of data, in order to extract the useful information, is a major undertaking and requires an automatic mechanism to aid with the extant repository of informa...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: International Journal of Advanced Computer Science and Applications

سال: 2022

ISSN: ['2158-107X', '2156-5570']

DOI: https://doi.org/10.14569/ijacsa.2022.0130962